Papers with Spearman’s Rank Correlation
Can Large Language Models Outperform Non-Experts in Poetry Evaluation? A Comparative Study Using the Consensual Assessment Technique (2025.emnlp-main)
Copied to clipboard
| Challenge: | Consensual Assessment Technique (CAT) for large language models is used to evaluate creativity, but is costly and time-consuming with non-experts. |
| Approach: | They adapt the Consensual Assessment Technique (CAT) for Large Language Models to a 90-poem dataset with a ground truth based on publication venue. |
| Outcome: | The proposed method outperforms the best human non-expert evaluations by significantly outperforming the best language models. |